Papers with modality-specific encoders

4 papers
Next Visit Diagnosis Prediction via Medical Code-Centric Multimodal Contrastive EHR Modelling with Hierarchical Regularisation (2024.findings-eacl)

Copied to clipboard

Challenge: Existing studies have not addressed the heterogeneous and hierarchical properties inherent in EHR data.
Approach: They propose a medical code-centric multimodal contrastive EHR learning framework with hierarchical regularisation that integrates multifaceted information encompassing medical codes, demographics, and clinical notes.
Outcome: The proposed framework integrates multifaceted information encompassing medical codes, demographics, and clinical notes using a tailored network design and bimodal contrastive losses.
LIFTED: Multimodal Clinical Trial Outcome Prediction via Large Language Models and Mixture-of-Experts (2025.findings-emnlp)

Copied to clipboard

Challenge: Clinical trials are costly and pivotal processes that require substantial expenses . a new approach to integrate multimodal data for clinical outcome prediction is needed .
Approach: a proposed framework transforms modality-specific data into natural language descriptions . a sparse Mixture-of-Experts mechanism then identifies shared patterns across modalities .
Outcome: a proposed framework outperforms baseline methods in predicting clinical trial outcomes . it transforms modality-specific data into natural language descriptions, encoded via unified encoders .
Beyond Logit Lens: Contextual Embeddings for Robust Hallucination Detection & Grounding in VLMs (2025.naacl-long)

Copied to clipboard

Challenge: Large Multimodal Models are plagued by hallucinations that limit their reliability and adoption.
Approach: They propose a method that leverages contextual token embeddings from LMMs to detect hallucinations.
Outcome: The proposed method improves hallucination detection and grounding across diverse categories while excelling in tasks requiring contextual understanding.
ImpliHateVid: A Benchmark Dataset and Two-stage Contrastive Learning Framework for Implicit Hate Speech Detection in Videos (2025.acl-long)

Copied to clipboard

Challenge: Existing studies on text-based hate speech detection focus on video-based approaches . however, hateful content remains a persistent challenge due to the vast amount of data generated every day.
Approach: They propose a novel two-stage contrastive learning framework for hate speech detection in videos . they train modality-specific encoders for audio, text, and image using contrastive loss .
Outcome: The proposed framework is based on two datasets, ImpliHateVid and HateMM datasets.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations